With more than 45 million cases awaiting disposal across Indian courts as of 2024, the judicial system faces an acute need for faster, smarter tools to support legal research. This work introduces an artificial-intelligence-driven legal research assistant tailored to the jurisprudence of the Supreme Court of India. Starting from a case description written in ordinary language, the system executes a four-stage pipeline: first, it performs dense semantic search by encoding the 26,688 judgments of the Indian Legal Documents Corpus (ILDC) with InLegalBERT and indexing roughly 1.1 million resulting text segments through a FAISS IVFFlat structure; second, it applies cross-encoder reranking to narrow the retrieved candidates down to the five precedents judged most semantically relevant; third, it forecasts the judicial outcome across three possible categories—Allowed, Dismissed, or Partly Allowed—via an InLegalBERT classification head paired with SHAP explainability; and fourth, it produces a structured reasoning summary in Issue–Rule–Application–Conclusion (IRAC) form using the Llama 3 8B model served locally through Ollama. By combining semantic search, neural reranking, interpretable outcome prediction, and AI-generated legal reasoning inside one coherent architecture, the framework helps legal practitioners locate relevant precedents, anticipate probable case outcomes, and streamline the overall research process. Constructed entirely from openly accessible Indian legal resources and transformer-based architectures, the system establishes an extensible base for intelligent legal support within the Indian judiciary.
Introduction
India's legal system faces a massive backlog of over 45 million pending cases (2024), making legal research increasingly difficult. Existing legal research tools, such as Indian Kanoon, rely primarily on keyword-based search, which often fails to identify legally similar cases expressed using different terminology. Additionally, these platforms generally lack outcome prediction, explainability, and structured legal reasoning. Although large language models like GPT-4 and Gemini have advanced natural language processing, they are not specialized for Indian law and may produce inaccurate or fabricated legal citations without retrieval-based grounding.
To address these challenges, this paper proposes an AI-powered legal research assistant for Indian Supreme Court cases that integrates multiple AI techniques into a single framework. The system uses InLegalBERT embeddings and a FAISS IVFFlat index over the Indian Legal Documents Corpus (ILDC) containing 26,688 Supreme Court judgments, divided into approximately 1.1 million text chunks. A cross-encoder reranks retrieved cases to identify the five most relevant precedents. An InLegalBERT-based classifier predicts whether a case is Allowed, Dismissed, or Partly Allowed, while SHAP explainability highlights the factors influencing the prediction. Finally, Llama 3 8B, running locally through Ollama, generates an Issue–Rule–Application–Conclusion (IRAC) legal reasoning summary grounded in the retrieved precedents.
The paper's key contributions include:
Semantic precedent retrieval using InLegalBERT and FAISS.
Cross-encoder reranking for improved retrieval accuracy.
Three-class judicial outcome prediction.
SHAP-based explainability for transparent predictions.
IRAC-style legal reasoning generation using an LLM.
Integration of all modules into a web-based legal research platform.
The literature review highlights progress in legal judgment prediction, domain-specific legal language models, dense retrieval, retrieval-augmented generation (RAG), and explainable AI. However, most existing systems address only individual tasks such as retrieval or prediction and do not provide a comprehensive solution for Indian legal research. The proposed framework fills this gap by combining retrieval, prediction, explainability, and reasoning within a single pipeline.
The methodology consists of five stages:
Data preprocessing: Clean and segment ILDC judgments into searchable text chunks.
Semantic retrieval: Generate InLegalBERT embeddings and retrieve relevant cases using FAISS.
Cross-encoder reranking: Improve ranking accuracy by jointly evaluating queries and candidate cases.
Outcome prediction: Classify cases into Allowed, Dismissed, or Partly Allowed.
Explainable reasoning: Use SHAP to explain predictions and Llama 3 8B to produce IRAC-formatted legal analysis.
The system is implemented as a modular web application with:
A data layer containing the indexed legal corpus.
A processing layer handling retrieval, reranking, prediction, explainability, and reasoning.
A FastAPI backend exposing REST APIs.
A React/Vite frontend for user interaction.
Evaluation uses standard information retrieval metrics (NDCG@5, MRR, Recall@20, Precision@5) and classification metrics (Accuracy, Precision, Recall, Macro-F1), while IRAC reasoning is assessed qualitatively. Experimental results indicate that semantic retrieval combined with cross-encoder reranking outperforms keyword-based search in identifying relevant precedents. The integrated pipeline provides transparent outcome predictions and structured legal reasoning, offering more comprehensive support than traditional legal research tools.
The study concludes that combining semantic retrieval, explainable AI, judicial outcome prediction, and retrieval-grounded LLM reasoning can significantly improve legal research for Indian Supreme Court cases. Current limitations include reliance on ILDC data only up to 2020, support only for English-language cases, and the need for improved factual consistency in generated reasoning. Future work will extend the corpus with newer judgments, add multilingual capabilities, further fine-tune retrieval models, and enhance the quality and reliability of AI-generated legal analysis.
Conclusion
This paper has described an AI-based legal research assistant for Indian Supreme Court jurisprudence that draws together semantic precedent retrieval, judicial outcome prediction, explainable AI, and structured legal reasoning within one framework. The system relies on InLegalBERT for dense retrieval and outcome classification, FAISS for fast similarity search, a cross-encoder to sharpen precedent relevance, SHAP to make predictions interpretable, and Llama 3 8B through Ollama to generate IRAC-formatted legal reasoning. It is built on the publicly available Indian Legal Documents Corpus (ILDC) [2], comprising 26,688 annotated Supreme Court judgments processed into approximately 1.1 million searchable text chunks.
By folding semantic retrieval, explainable outcome prediction, and AI-assisted reasoning into a single application, the framework helps legal professionals find relevant precedents, gauge likely outcomes, and obtain structured analysis more efficiently than conventional keyword-based research tools allow. Its modular design also means individual components can be upgraded independently, keeping the system adaptable as legal AI continues to develop.
Looking ahead, future work will aim to strengthen retrieval through domain-specific fine-tuning, extend the corpus to more recent judgments and additional Indian courts, add support for multilingual documents and queries, and continue improving the quality and factual reliability of the AI-generated reasoning. These directions are expected to broaden the practical usefulness of the system and contribute more generally to the advancement of intelligent legal research in India.
References
[1] National Judicial Data Grid (NJDG), Government of India, \"Pendency Statistics,\" Ministry of Law and Justice, New Delhi, 2024. [Online]. Available: https://njdg.ecourts.gov.in/
[2] S. K. Malik, A. Guha, P. Pandey, A. Ghosh, S. Ghosh, and P. Majumder, \"ILDC for CJPE: Indian legal documents corpus for court judgment prediction and explanation,\" in Proc. 59th Annual Meeting of the ACL and 11th Int. Joint Conf. on NLP (ACL-IJCNLP), pp. 4046–4062, Aug. 2021.
[3] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, \"BERT: Pre-training of deep bidirectional transformers for language understanding,\" in Proc. NAACL-HLT, pp. 4171–4186, 2019.
[4] I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, \"LEGAL-BERT: The muppets straight out of law school,\" in Proc. Findings of EMNLP, pp. 2898–2904, 2020.
[5] N. Reimers and I. Gurevych, \"Sentence-BERT: Sentence embeddings using Siamese BERT-networks,\" in Proc. EMNLP-IJCNLP, pp. 3982– 3992, 2019.
[6] V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, \"Dense passage retrieval for open-domain question answering,\" in Proc. EMNLP, pp. 6769–6781, 2020.
[7] J. Johnson, M. Douze, and H. Jégou, \"Billion-scale similarity search with GPUs,\" IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, 2021.
[8] R. Nogueira and K. Cho, \"Passage re-ranking with BERT,\" arXiv preprint arXiv:1901.04085, 2019.
[9] G. Rosa, R. Rodrigues, A. de Alencar Lotufo, and R. Nogueira, \"Yes, BM25 is a strong baseline for legal case retrieval,\" in Proc. SIGIR Workshop on Legal Information Retrieval, 2022.
[10] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, \"Retrieval-augmented generation for knowledge-intensive NLP tasks,\" in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 9459–9474, 2020.
[11] S. M. Lundberg and S.-I. Lee, \"A unified approach to interpreting model predictions,\" in Advances in Neural Information Processing Systems (NIPS), vol. 30, pp. 4765–4774, 2017.
[12] S. Jain and B. C. Wallace, \"Attention is not explanation,\" in Proc. NAACL-HLT, pp. 3543–3556, 2019.
[13] N. Aletras, D. Tsarapatsanis, D. Preo?iuc-Pietro, and V. Lampos, \"Predicting judicial decisions of the European Court of Human Rights: A natural language processing perspective,\" PeerJComput. Sci., vol. 2, e93, 2016.
[14] D. M. Katz, M. J. Bommarito, and J. Blackman, \"A general approach for predicting the behavior of the Supreme Court of the United States,\" PLOS ONE, vol. 12, no. 4, e0174698, 2017.
[15] H. Zhong, Z. Guo, C. Tu, C. Xiao, Z. Liu, and M. Sun, \"Legal judgment prediction via topological learning,\" in Proc. EMNLP, pp. 3540– 3549, 2018.
[16] I. Chalkidis, A. Jana, D. Hartung, M. Bommarito, I. Androutsopoulos, D. M. Katz, and N. Aletras, \"LexGLUE: A benchmark dataset for legal language understanding in English,\" in Proc. ACL, pp. 4310–4330, 2022.
[17] Y. Feng, C. Li, and N. Ng, \"Legal case retrieval: A survey of the state of the art,\" in Proc. 45th Int. ACM SIGIR Conf. on Research and Development in Information Retrieval, pp. 2937–2946, 2022.
[18] H. Liao, C. Qin, Y. Ren, H. Li, Z. Huang, Y. Zhang, and C. Wang, \"VERDICT: Verifiable evolving reasoning with directive-informed collegial teams for legal judgment prediction,\" arXiv preprint arXiv:2603.19306, 2026.
[19] S. K. Nigam, B. D. Patnaik, S. Mishra, A. V. Thomas, N. Shallum, K. Ghosh, and A. Bhattacharya, \"NyayaRAG: Realistic legal judgment prediction with RAG under the Indian common law system,\" in Proc. 14th IJCNLP and 4th Conf. of the Asia-Pacific Chapter of the ACL (IJCNLP-AACL), pp. 1709–1726, 2025.
[20] Y. Zhang, Z. Tian, S. Zhou, H. Wang, W. Hou, Y. Liu, and B. Zhou, \"RLJP: Legal judgment prediction via first-order logic rule-enhanced with large language models,\" arXiv preprint arXiv:2505.21281, 2025.
[21] S. K. Nigam, A. Deroy, S. Maity, and A. Bhattacharya, \"Rethinking legal judgment prediction in a realistic scenario in the era of large language models,\" in Proc. Natural Legal Language Processing Workshop 2024 (NLLP), pp. 61–80, 2024.
[22] S. K. Nigam and A. Deroy, \"Fact-based court judgment prediction,\" in Proc. 15th Annual Meeting of the Forum for Information Retrieval Evaluation (FIRE), pp. 78–82, 2023.
[23] J. Cui, Z. Li, Y. Yan, B. Chen, and L. Yuan, \"ChatLaw: A multi-agent collaborative legal assistant with knowledge graph enhanced mixtureof-experts large language model,\" arXiv preprint arXiv:2306.16092, 2023.
[24] N. Guha, J. Nyarko, D. E. Ho, C. Ré, A. Chilton, A. Chohlas-Wood, A. Peters, B. Waldon, D. N. Rockmore, and D. Zambrano, \"LegalBench: A collaboratively built benchmark for measuring legal reasoning in large language models,\" in Advances in Neural Information Processing Systems (NeurIPS), 2024.
[25] A. Blair-Stanek, N. Holzenberger, and B. Van Durme, \"Can GPT-3 perform statutory reasoning?\" in Proc. Natural Legal Language Processing Workshop (NLLP), pp. 213–224, 2023.
[26] H. Touvron, T. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, et al., \"Llama 2: Open foundation and fine-tuned chat models,\" arXiv preprint arXiv:2307.09288, 2023.
[27] W. Yang, W. Jia, X. Zhou, and Y. Luo, \"Legal judgment prediction via multi-perspective bi-feedback network,\" in Proc. 28th Int. Joint Conf. on Artificial Intelligence (IJCAI), pp. 4085–4091, 2019.
[28] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, and D. Zhou, \"Chain-of-thought prompting elicits reasoning in large language models,\" in Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 24824–24837, 2022.